Skip to content

Run reports and the behaviour check on a local model - #110

Open
alanhc wants to merge 6 commits into
sysprog21:mainfrom
alanhc:shim-tools
Open

alanhc wants to merge 6 commits into
sysprog21:mainfrom
alanhc:shim-tools

Conversation

@alanhc

@alanhc alanhc commented Sep 25, 2026 •

Copy link
Copy Markdown
Collaborator

This now carries #99 as well, as asked there. Main had moved a long way since both branches were cut (multi-key reports, doubling backoffs, report recovery, the played-candidate check from #107), so rather than replay 16 commits over it, the work is rebuilt on current main as five commits. The old history is still readable on #99.

Summary

The report, the quiet-pause reviews and the interviewer behaviour check can now run on a model on the operator's own GPU instead of Gemini. The live interviewer is untouched and still talks to Gemini.

  • CODETRIAL_GEMINI_REST_BASE, read from the environment like INTERVIEW_ROOM_NAME, points every generateContent call at another server. Unset, the URL is what it was.
  • scripts/gemini-shim.py (standard library only) answers generateContent from llama-server's OpenAI-compatible endpoint. It converts the response schema to the JSON Schema llama.cpp compiles to a grammar, carries function declarations and calls both ways, passes upstream status codes through so a 503 still reads as a 503, and gives up on llama-server when the caller would. A request with no output limit gets one, and --thinking off turns thinking off for every request, which Gemma 4 needs for the behaviour check. tests/test_gemini_shim.py runs in the gate.
  • Longer deadlines for a local base only. Any base other than Google's gets 45 seconds a call and 250 for the report, against 20 and 125; 250 is what five calls and the doubling backoffs need. Only the server knows which is in force, so the page's wait before offering to leave and its wait on a regenerated report now come from /runtime-config.js, and the page's own values stay the hosted ones as floors.
  • Repairs that say what to fix. When a report names the published problem, the repair names the title and the scenario name to use; when the plan and the feedback disagree, it names each wrong item by index, each improvement with no item, and the swap when there is one of each. Told only that a rule failed, the local model sent back the same output byte for byte. None of this reaches the candidate's failure note.
  • The behaviour check follows the base, so pointed at the shim it sends what production sends. BEHAVIOR_LOCAL_BASE from Hold played candidates to the interview rules #107 stays as the direct route, without the shim.

Numbers

RTX 5070 Ti, gemma-4-12b Q4_K_M on llama.cpp, shim with --thinking off:

  • Reports on the Two Sum golden prompt, this branch: 4 of 4 valid, 14 to 16 seconds each. On the old branches, the plan repair took this from 0 of 6 to 10 of 10, and the title repair from 2 of 4 to 4 of 4.
  • Behaviour check through the shim: the Two Sum script passed in 9 seconds. On the old branch, 18 of 21 problem runs across seven runs; each miss was a second hint request answered without calling log_hint.

One prompt, one machine: read these as "this works end to end", not as a comparison with Gemini.

Test plan

  • The report budget test holds both deadline pairs to five calls plus backoffs; the escape-hatch and recovery-limit tests hold the page's floors to the hosted deadline and the server's values to both.
  • binary_web_gives_a_local_report_base_the_longer_wait starts the binary with the base set and checks /runtime-config.js says 260000 and 265; without it, the runtime-config test pins 135000 and 140.
  • report-recovery.test.js: the server can raise the retry wait and cannot lower it.
  • Repair guidance: the title test, and seven plan tests (reworded, repeated, several at once, counted once across sections, a missing weakness keeping indexes, a matching plan getting none, guidance staying out of the failure note).
  • a_local_model_writes_a_report is an ignored unit test that makes one real report through the base.
  • ./scripts/test.sh passes locally: 844 JS and 1228 Rust tests, none skipped, formatting clean. Every commit builds clean under clippy. cargo mutants --in-diff on this diff: 42 mutants, 38 caught, 4 unviable, none missed. cargo-audit and shellcheck were not installed here.

Not tested: a full interview end to end with the report coming from a local model, and anything against Gemini.

@alanhc alanhc changed the title Let the behaviour check run against a local model Run reports and the behaviour check on a local model Oct 3, 2026
@alanhc
alanhc marked this pull request as ready for review October 3, 2026 06:21

@cubic-dev-ai cubic-dev-ai Bot left a comment •

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

All reported issues were addressed across 19 files

Reply with feedback, questions, or to request a fix.

Re-trigger cubic

Comment thread tests/web/assets.rs Outdated
Comment thread scripts/gemini-shim.py
Comment thread web/interview.js Outdated
Comment thread web/report-recovery.js
Comment thread scripts/gemini-shim.py
Comment thread README.md Outdated
Comment thread src/gemini.rs
alanhc added 5 commits October 7, 2026 00:42
The report and interim-review calls always went to Google, so trying
a model on one's own GPU meant editing the URL by hand. They now go to
CODETRIAL_GEMINI_REST_BASE when it is set, read from the process
environment like INTERVIEW_ROOM_NAME; unset, nothing changes. The live
socket is untouched.

scripts/gemini-shim.py answers generateContent from llama-server's
OpenAI-compatible endpoint. It turns the Gemini schema into the JSON
Schema llama.cpp compiles to a grammar, carries function declarations
and calls both ways, passes upstream status codes through so the retry
rules still see a 503 as a 503, and gives up on llama-server when the
caller would. It caps a request that names no output limit, so a model
stuck repeating itself ends as MAX_TOKENS instead of out-waiting the
shim, and --thinking off turns thinking off for every request, which
Gemma 4 needs for a check that names no thinking budget.
tests/test_gemini_shim.py covers the mapping and runs in the gate.

An ignored unit test makes one real report through the base. Against
gemma-4-12b on an RTX 5070 Ti it wrote a valid report in 16 seconds.
A 9B to 14B model on one consumer GPU takes 14 to 32 seconds per
report call, against a 20-second attempt limit sized for Gemini, so a
local model timed out on most calls. When CODETRIAL_GEMINI_REST_BASE
points anywhere but Google, the attempt limit is 45 seconds and the
report deadline 250, so five attempts and their doubling backoffs
still fit; Gemini keeps 20 and 125. Only the server knows which
deadline is in force, so the page's wait before offering to leave and
its wait on a regenerated report now come from /runtime-config.js, and
the page's own values stay the hosted ones as floors.
A report whose improvement plan called the exercise "a 'Two Sum'
style problem" was sent back with only "names the published problem"
and the field's path. The model could not tell which words broke the
rule, wrote the same sentence twice more, and the report was lost.

The repair prompt now names the published title and the scenario's
title to use instead. The title stays out of the error itself: that
error is the failure note the candidate reads, and the title is what
it must not show them. The original prompt already carries the title,
so the model learns nothing new from it.

Against gemma-4-12b on the Two Sum weak-candidate prompt, 2 of 4
reports passed before this and 4 of 4 after, three of them on the
first repair.
When the improvement plan and the feedback disagreed, the repair was
told only that they did, and a local 12B model sent back the same plan
byte for byte on every repair: told that something in a list of four
was wrong, it could not find which. The usual cause is a weakness
reworded on its way into the plan, or a repeat standing where an
improvement was left out. The repair now names each wrong item by
index and each improvement with no item, and when there is one of
each, the swap. Improvements are counted once each, the way the
validator counts them, and an item without a weakness keeps its index.
The note the candidate reads is unchanged.

Against gemma-4-12b on the Two Sum report prompt, 0 of 6 reports
passed before this and 10 of 10 after; 7 of the 10 broke the plan on
their first attempt and were repaired.
The scripted check posted to Google's URL by hand, so it could not
follow CODETRIAL_GEMINI_REST_BASE the way the report does. It now
builds its URL with gemini_generate_content_url, so pointed at
scripts/gemini-shim.py it sends what production sends through the
shim the report uses, tools included. The direct route the played
candidates added, BEHAVIOR_LOCAL_BASE, stays for talking to an
OpenAI-compatible server without the shim. README and
docs/development.md describe both. Against gemma-4-12b through the
shim with --thinking off, the Two Sum script passed in 9 seconds.
@alanhc

alanhc commented Oct 6, 2026

Copy link
Copy Markdown
Collaborator Author

I ran one interview by hand with everything local: two-sum (Chargeback Pair Match), Coding + behavioral, Editor, Python, about five minutes, with me speaking into the microphone. Nothing in the session reached Google.

Setup:

Server Intel Core i7-14700 (20 cores, 28 threads), 64 GB RAM, RTX 5070 Ti 16 GB (driver 580.173), Ubuntu 24.04, kernel 7.0
Client MacBook Pro 14" (M1 Pro), macOS 26.6.2 (25G83), Chrome 154.0.8037.98 (arm64), built-in camera and microphone, connected to the server over Tailscale with an SSH tunnel for ports 8095, 7880 and 7881, so the page was on http://127.0.0.1:8095 and the microphone was allowed
CodeTrial this branch at ed418ae, plus one local commit that adds CODETRIAL_GEMINI_LIVE_URL, which is not part of this PR
Reports and text calls scripts/gemini-shim.py --thinking off in front of llama.cpp llama-server (bd43117) with gemma-4-12b-it Q4_K_M, -ngl 99 -c 32768 -fa on -ctk q8_0 -ctv q8_0 --jinja
Live interviewer a local speech pipeline behind CODETRIAL_GEMINI_LIVE_URL: Silero VAD, faster-whisper 1.2.1 large-v3-turbo, the same gemma, Kokoro 1.0 on onnxruntime-gpu 1.30
LiveKit self-hosted livekit-server 1.13.8 in --dev mode on the server
GPU memory 15.0 of 16.3 GB in use: gemma 8.2, speech models 2.4, and 4.0 for an unrelated process

What this PR covers:

  • The report came back on the first call in 16.6 s (4,533 prompt tokens, 1,215 out), with no repair, and was delivered on the first attempt (report_delivery ... attempt=1 outcome=acknowledged). That is well inside the 45 s per-call limit and the 250 s local deadline.
  • The hint count was right: 2. A third request was withheld and not counted, as record_hint intends.
  • The summary called the problem by its scenario name and never by the LeetCode title. The behavioral round showed as skipped, which it was, since I ended at about five minutes.
  • The other 16 text calls (15 phase checks and 1 pause review) also went through the shim and all finished, each in 0.7 to 3.6 s.
  • One mistake in the report: it said my loop inserted before checking. My code actually inserted only inside the if, so it never stored anything. Its other coding point, that the function name did not match the starter, was right. I had typed the wrong name, so the tests never ran.

Found on the interviewer side, outside this PR:

  • gemma never called record_framework_evidence, so no Algorithm evidence was recorded and the third hint was withheld even after I had described the approach.
  • whisper made up a sentence I did not say ("I'm going to go to the next slide."), and the interviewer answered it.
  • After I gave the complexity, the interviewer asked for it twice more.

This is one run on one problem. I will run a few more, including a whiteboard interview, before calling the local report path done.

The phase judge asks for application/json with no response schema, and
the shim passed that on unconstrained. gemma-4-12b then wrapped its
answer in a ```json fence, the judge's parser refused it, and in a
five-minute interview by hand all 15 judgments were dropped: no REACTO
step was recorded from speech, so the last hint stayed withheld after
the approach had been stated. Such a request now asks llama-server for
a JSON object, which is what Gemini returns for it.

The live judge test now follows CODETRIAL_GEMINI_REST_BASE, as the rest
of the behaviour check does, and prints each judgment. Against the shim
it failed before this change and passes after it.
@alanhc

alanhc commented Oct 6, 2026

Copy link
Copy Markdown
Collaborator Author

The first run found a bug in this PR, now fixed in 4d63990, and a second run with the same setup confirms the fix.

The bug. The phase judge requests application/json without a response schema, and scripts/gemini-shim.py passed that through to llama-server with no constraint. gemma-4-12b wrapped its answer in a ```json fence, apply_phase_judgment could not parse it, and all 15 judgments in the first run were dropped. No REACTO step was recorded from speech, and because no Algorithm evidence existed, the last hint stayed withheld even after I had described my approach. The shim now asks llama-server for a JSON object whenever the request asks for JSON, which matches what Gemini returns. the_phase_judge_ticks_what_the_candidate_said_and_nothing_else in tests/interview_behavior.rs now follows CODETRIAL_GEMINI_REST_BASE and prints each judgment. Against the shim, it failed before the fix and passes after it.

Second run, same problem and setup, about seven minutes:

Run 1 Run 2
REACTO steps recorded 0 all 6: five first recorded by the phase judge and Test by the test run; the interviewer also recorded Coding and Optimizations again
Hints counted 2 (third withheld) 3 (third given after the approach was stated)
Phase-judge calls 15, all dropped 11
Report 16.6 s, first call, delivered on attempt 1 17.6 s (5,255 prompt / 1,236 out), first call, delivered on attempt 1
Tests never ran (I typed the wrong function name) failed on [3,2,4], 6 as intended, then 5/5 plus my case after the fix

The report this time described the bug correctly (inserting before checking, so 3 matched itself), scored all six REACTO phases, marked the behavioral round as skipped, and did not name the LeetCode title.

What the judge still gets wrong with gemma. The format is fixed; some of its judgments are not good:

  • It recorded Optimizations at 5:16, quoting my brute-force line ("check every pair with two nested loops which is O of n squared"). I gave the actual time and space trade-off at 6:00. The quote passes quote_is_grounded because those are my words, but they do not show that step.
  • It recorded Algorithm with the same brute-force line, not the one-pass dictionary plan I stated later.
  • The interviewer's own two record_framework_evidence calls used a confidence of 5. That field is 0 to 100, so gemma appears to be treating it as a fraction.

These are judgment errors by a 12B model, not transport problems, so I have not tried to fix them in the shim. If the check should also hold the quote to the phase it claims, that would belong in apply_phase_judgment rather than in this PR.

@alanhc

alanhc commented Oct 6, 2026

Copy link
Copy Markdown
Collaborator Author

A correction to my first comment, where I wrote "Nothing in the session reached Google". That was not quite true.

  • LiveKit STUN, both runs. livekit-server --dev with no rtc.stun_servers set gives every participant stun.l.google.com, stun1.l.google.com and global.stun.twilio.com as ICE servers. So the browser and the agent most likely sent STUN binding requests to Google and Twilio. These requests only ask for the public address; no audio, transcript or code goes with them. Media went over 127.0.0.1:7881 through the tunnel.
  • Hugging Face, run 1 only. faster-whisper checked the Hugging Face hub when it loaded its cached model. I set HF_HUB_OFFLINE=1 before run 2.

Every model call was local in both runs: the report, the phase judge and the pause review went to gemma through the shim, and the live interviewer used local whisper, gemma and Kokoro. Neither CodeTrial's log nor either shim's log shows a request to googleapis.com. LiveKit now runs with rtc.stun_servers: [127.0.0.1:3478], and a test join returned only that server, so later runs will make no STUN requests to Google or Twilio.

@alanhc

alanhc commented Oct 7, 2026

Copy link
Copy Markdown
Collaborator Author

I opened #257 to collect results from other models and GPUs. It gives step-by-step instructions and a script that runs the report test three times plus the behaviour check against this branch, along with a results template. My two runs on the 5070 Ti are in it as the reference. The results should show whether the fixed local deadlines (45 s per report call, 12 s per side call) hold on slower hardware, and which models follow the hint and disclosure rules.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant